Papers with automated evaluation methods
TempViz: On the Evaluation of Temporal Knowledge in Text-to-Image Models (2026.eacl-long)
Copied to clipboard
| Challenge: | Existing studies on temporal knowledge in text-to-image models have not explored how temporal phenomena are handled in text models. |
| Approach: | They propose a data set to holistically evaluate temporal knowledge in image generation using 7.9k prompts and more than 600 reference images. |
| Outcome: | The proposed model evaluates temporal knowledge in image generation using 7.9k prompts and more than 600 reference images. |
ConQRet: A New Benchmark for Fine-Grained Automatic Evaluation of Retrieval Augmented Computational Argumentation (2025.naacl-long)
Copied to clipboard
| Challenge: | Existing methods for evaluating RAArg are costly and lack long, complex arguments and real-world evidence. |
| Approach: | They propose to use multiple fine-grained LLM judges to evaluate RAArg using a new benchmark that features long and complex human-authored arguments on debated topics. |
| Outcome: | The proposed methods provide better and more interpretable assessments than traditional single-score metrics and even previously reported human crowdsourcing. |
LitBench: A Benchmark and Dataset for Reliable Evaluation of Creative Writing (2026.eacl-long)
Copied to clipboard
| Challenge: | a single prompt can inspire countless valid stories, making objective verification impossible. |
| Approach: | They propose a large-scale benchmark for creative writing evaluation using a reddit corpus and a 2,480-pair test set. |
| Outcome: | The proposed model outperforms existing OTS judges and generative reward models in the evaluation of creative writing. |
Automated Metrics for Medical Multi-Document Summarization Disagree with Human Evaluations (2023.acl-long)
Copied to clipboard
Lucy Lu Wang, Yulia Otmakhova, Jay DeYoung, Thinh Hung Truong, Bailey Kuehl, Erin Bransom, Byron Wallace
| Challenge: | Prior work has shown that models may exploit shortcuts that are difficult to detect using standard n-gram similarity metrics such as ROUGE. |
| Approach: | They propose to use human-assessed summary quality facets and pairwise preferences to improve MDS evaluation methods. |
| Outcome: | The proposed methods improve the quality of literature review summarization models . they use human-assessed summary quality facets and pairwise preferences . |
Evaluating Saliency Explanations in NLP by Crowdsourcing (2024.lrec-main)
Copied to clipboard
| Challenge: | a crowdsourced method to evaluate saliency methods in NLP is proposed . saliencies are difficult for humans to understand, and can cause psychological harm . |
| Approach: | They propose a method to evaluate saliency methods in NLP by crowdsourcing . they recruited 800 crowd workers and empirically evaluated seven salience methods . |
| Outcome: | The proposed method evaluates saliency methods on two datasets using crowdsourced data . it shows that the results are comparable to existing methods on NLP and CV fields . |
AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific Research (2025.acl-long)
Copied to clipboard
Yilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan
| Challenge: | a benchmark designed to evaluate the capabilities of LLMs in designing ablation studies for scientific research is available online. |
| Approach: | They propose to use a benchmark to evaluate LLMs' ability to design ablation studies . they investigate whether current automated evaluation methods are not reliable . |
| Outcome: | The benchmark compared leading LLMs with human experts on generating detailed ablation study designs . the results show that current evaluation methods are not reliable for the task . |
Revisiting Automated Evaluation for Long-form Table Question Answering (2024.emnlp-main)
Copied to clipboard
| Challenge: | Existing automated metrics for long-form table question answering (LFTQA) are poorly correlated with human judgments and fail to distinguish between factually accurate responses and those that are factual incorrect. |
| Approach: | They propose to use a meta-evaluation dataset to assess the effectiveness of LLM-based LFTQA systems. |
| Outcome: | The proposed meta-evaluation dataset includes 2,988 human-annotated examples. |
Igniting Creative Writing in Small Language Models: LLM-as-a-Judge versus Multi-Agent Refined Rewards (2025.emnlp-main)
Copied to clipboard
| Challenge: | Existing methods for enhancing Large Language Models (LLMs) struggle with novelty and Reinforcement Learning from human feedback (RLHF) is costly. |
| Approach: | They propose to use a Reward Model (RM) and a principle-guided LLM-as-a-Judge to enhance creative output over baselines. |
| Outcome: | The proposed approach significantly enhances creative output over baselines, but the principle-guided LLM-as-a-Judge yields superior generation quality. |
Hi Guys or Hi Folks? Benchmarking Gender-Neutral Machine Translation with the GeNTE Corpus (2023.emnlp-main)
Copied to clipboard
| Challenge: | Societal gender asymmetries and inequalities are perpetuated through language . MT often defaults to masculine representations by making undue binary gender assumptions . |
| Approach: | They propose a benchmark and automated evaluation methods to assess gender-neutral translation from English to Italian. |
| Outcome: | The proposed method is based on a survey on gender-neutral translation. |